Repository navigation
[Model][NVIDIA] Route DSA models to the CUDA non-compiled path - #52861
Conversation
Route GlmMoeDsaForCausalLM and the matching MTP architecture to the optimized deepseek_v32 implementation on SM100-family devices, keeping the generic deepseek_v2 fallback everywhere else. Default the KV cache to FP8 since this implementation requires a sparse FP8 cache. Split out of vllm-project#48597 (reverted by vllm-project#49768). Co-authored-by: Claude <noreply@anthropic.com> Signed-off-by: Peiyuan Zhou <peiyuanzhou1994@gmail.com>
mypy rejects rebinding a class name with a plain assignment ("Cannot assign to
a type"), which the other branches of this dispatch bind by import. Use an
import alias so every branch defines the name the same way.
Co-authored-by: Claude Opus 5
Signed-off-by: Peiyuan Zhou <peiyuanzhou1994@gmail.com>
Routing GlmMoeDsaForCausalLM here changed what a default launch does. This implementation asserted an fp8 KV cache, so an unset --kv-cache-dtype was rewritten to fp8 from inside a per-layer constructor: the engine went on reporting kv_cache_dtype=auto while allocating an fp8 cache (2,935,232 KV tokens against main's 1,509,888 for the same launch), and the "Using ... data type to store kv cache" line never printed. The assert was stricter than anything below it needs. fused_norm_rope already writes an unquantized cache when the dtype is not fp8, fused_q already emits the bf16 (ql_nope, q_pe) query that the fp8_ds_mla layout uses, FlashInfer sparse accepts that tuple, and the ROCm subclass already derives the same two flags from the cache dtype. Only the query form (taken from the backend capability rather than the cache) and the unconditional fp8 view of the paged cache assumed fp8; both now come from the dtype, and nothing rewrites cache_dtype. Measured on 8xB300, GLM-5.2 block-fp8, TP8, MTP=5, 8192-in/1024-out at concurrency 1. Default launch: 1,513,024 KV tokens (bf16, matching main's 1,509,888), 362 tok/s vs main's 340 for the same config, GSM8K 0.950. --kv-cache-dtype fp8_e4m3 is unchanged at 2,935,232 tokens and 403 tok/s, GSM8K 0.938. Co-authored-by: Claude Opus 5 Signed-off-by: Peiyuan Zhou <peiyuanzhou1994@gmail.com>
Choose Model Runner V2 for GLM-5.2 and auto-enable the non-compiled breakable CUDA graph path for the model and MTP architectures on every platform. The optimized SM100 model routing remains hardware-specific, while this serving default does not. Cover both NVFP4 and FP8 checkpoints with and without MTP. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
GLM-5.2 now defaults to the non-compiled MRV2 path, so route every CUDA device through the NVIDIA deepseek_v32 implementation. Capability-specific kernels continue to gate themselves and fall back when unavailable. Drop the standalone routing and KV-cache-form test files. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Use the deepseek_v32 main and MTP implementations for DeepSeek V3.2, including NVIDIA NVFP4 checkpoints. Default both architectures to MRV2 with breakable CUDA graphs and remove the obsolete MTP eager override. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Resolve the registry through CUDA-specific aliases so NVIDIA uses the deepseek_v32 implementation while ROCm, XPU, and CPU retain their existing defaults. Keep the explicit AMD package exports available for opt-in use. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Move the CUDA-or-generic registry dispatch into a separate module that exports DeepseekV32ForCausalLM and DeepseekV32MTP without prefixed aliases. Keep ROCm on the generic compiled MRV1 path without automatic breakable graphs while preserving the opt-in AMD package exports. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
/ci run |
|
✅ Triggered Buildkite CI #84554 for commit |
ResultsNVFP4 TP4, dummy weights, no MTPThis is pure tensor parallelism (
Candidate mean TPOT was also lower at every point: 7.41/8.76/15.51/43.34/ NVFP4 TP4, real weights, MTP3This is also pure tensor parallelism (DP=1, EP disabled). Both variants loaded
Candidate mean TPOT was lower at every point: 2.94/6.57/14.32/37.60/39.27 ms NVFP4 DEP4, dummy weights, no MTPThese runs use TP1 x DP4 with expert parallelism enabled. They are therefore
Candidate had more KV capacity per replica (367,616 versus 284,544 tokens) NVFP4 DEP4, real weights, MTP3These runs use TP1 x DP4 with expert parallelism enabled and are not subject
Candidate mean TPOT is lower at concurrency 256/1024: 12.65/12.59 ms versus FP8 TP8, dummy weights, no MTPThis is pure tensor parallelism (
The candidate is faster at every measured concurrency and every point has zero FP8 TP8, real weights, MTP3This is pure tensor parallelism (DP=1, EP disabled). Both variants loaded the
Candidate mean TPOT is lower at every point: 3.67/8.50/17.32/54.39/55.59 ms FP8 DEP8, dummy weights, no MTPThese runs use TP1 x DP8 with expert parallelism enabled, four local DP ranks
Candidate KV capacity is 461,440 tokens per replica versus 336,064 for the FP8 DEP8, real weights, MTP3These runs use TP1 x DP8 with expert parallelism enabled, four local DP ranks
Both variants completed every request. Candidate mean TPOT is lower at every Real-weight output-quality checkA paired non-speculative TP4 evaluation used the real NVFP4 checkpoint, MRV2,
The candidate has no accuracy regression on this fixed sample. Its single |
…project#52861) Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com>
### What this PR does / why we need it? The PR adapts vllm-ascend for compatibility with the latest vLLM main (commit `ba07e4a4`). | Files | Upstream vLLM change | vllm-ascend adaptation | |-------|---------------------|------------------------| | `.github/vllm-main-verified.commit` | — | Updated verified main commit hash from `cdc4824a21` to `ba07e4a48` | | `tests/e2e/conftest.py` | [vllm#53272](vllm-project/vllm#53272) — upstream plans to remove native Hunyuan V1/VL; [vllm#51665](vllm-project/vllm#51665) — dropped HunYuanVL `lm_head` workaround | Added `skip` condition for HunyuanVL e2e when `vllm_version_is("0.27.1")` is False (vLLM main) | | `tests/ut/core/test_profiling_chunk.py` | — vLLM main `Scheduler.__init__` reads `model_config.uses_mrope`, which infinitely recurses on a bare MagicMock | Version-gated: sets `type(model_config).uses_mrope = PropertyMock(return_value=False)` on main | | `tests/ut/core/test_recompute_scheduler.py` | — vLLM main `Scheduler.add_request` reads `spec_decode_metrics_level` | Version-gated: sets `scheduler.spec_decode_metrics_level = "none"` on main | | `tests/ut/patch/platform/test_patch_structured_output.py` | — Upstream changed structured output validation error type from `ValueError` to `VLLMValidationError` | Version-gated: `error_type = ValueError if vllm_version_is("0.27.1") else VLLMValidationError` used in all three fake validation functions and `pytest.raises` | | `vllm_ascend/attention/attention_v1.py` | [vllm#52839](vllm-project/vllm#52839) — moved `pcp.py` from `vllm.model_executor.layers.attention.pcp` to `vllm.v1.attention.ops.pcp` | Version-gated `_gather_prefill_cache_inputs` import path | | `vllm_ascend/compilation/acl_graph.py` | [vllm#49134](vllm-project/vllm#49134) — `get_current_vllm_config()` now raises `AssertionError` when called outside `set_current_vllm_config()` context | `update_full_graph_params` version-gated: main branch wraps `get_impl_cls()`/`update_graph_params` in `with set_current_vllm_config(vllm_config):` | | `vllm_ascend/models/deepseek_mtp.py` | [vllm#53106](vllm-project/vllm#53106) — removed `skip_prefixes` kwarg from `AutoWeightsLoader.__init__` | `AscendGlmMoeDsaForCausalLM.load_weights` version-gated: 0.27.1 uses `skip_prefixes=["rot."]`; main uses `WeightsMapper(orig_to_new_prefix={"rot.": None})` passed via `load_weights(weights, mapper=mapper)` | | `vllm_ascend/spec_decode/llm_base_proposer.py` | [vllm#52861](vllm-project/vllm#52861) — added `DeepseekV32MTPModel` to MTP architecture set in `model_returns_tuple()` | `model_returns_tuple` version-gated: 0.27.1 checks `{"DeepSeekMTPModel", "KimiK3MTPModel"}`; main also includes `"DeepseekV32MTPModel"` | | `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` | [vllm#52188](vllm-project/vllm#52188) — added `cp_rank`, `CP_SIZE`, `CP_INTERLEAVE` params to `_prepare_dflash_inputs_kernel` for DCP support | Entire `_prepare_dflash_inputs_kernel_ascend` kernel duplicated under `vllm_version_is("0.27.1")` gate: 0.27.1 uses 30 pos + 3 constexpr (no DCP params); main uses 31 pos + 5 constexpr | ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@cdc4824 --------- Signed-off-by: hfadzxy <starmoon_zhang@163.com>
### What this PR does / why we need it? The PR adapts vllm-ascend for compatibility with the latest vLLM main (commit `ba07e4a4`). | Files | Upstream vLLM change | vllm-ascend adaptation | |-------|---------------------|------------------------| | `.github/vllm-main-verified.commit` | — | Updated verified main commit hash from `cdc4824a21` to `ba07e4a48` | | `tests/e2e/conftest.py` | [vllm#53272](vllm-project/vllm#53272) — upstream plans to remove native Hunyuan V1/VL; [vllm#51665](vllm-project/vllm#51665) — dropped HunYuanVL `lm_head` workaround | Added `skip` condition for HunyuanVL e2e when `vllm_version_is("0.27.1")` is False (vLLM main) | | `tests/ut/core/test_profiling_chunk.py` | — vLLM main `Scheduler.__init__` reads `model_config.uses_mrope`, which infinitely recurses on a bare MagicMock | Version-gated: sets `type(model_config).uses_mrope = PropertyMock(return_value=False)` on main | | `tests/ut/core/test_recompute_scheduler.py` | — vLLM main `Scheduler.add_request` reads `spec_decode_metrics_level` | Version-gated: sets `scheduler.spec_decode_metrics_level = "none"` on main | | `tests/ut/patch/platform/test_patch_structured_output.py` | — Upstream changed structured output validation error type from `ValueError` to `VLLMValidationError` | Version-gated: `error_type = ValueError if vllm_version_is("0.27.1") else VLLMValidationError` used in all three fake validation functions and `pytest.raises` | | `vllm_ascend/attention/attention_v1.py` | [vllm#52839](vllm-project/vllm#52839) — moved `pcp.py` from `vllm.model_executor.layers.attention.pcp` to `vllm.v1.attention.ops.pcp` | Version-gated `_gather_prefill_cache_inputs` import path | | `vllm_ascend/compilation/acl_graph.py` | [vllm#49134](vllm-project/vllm#49134) — `get_current_vllm_config()` now raises `AssertionError` when called outside `set_current_vllm_config()` context | `update_full_graph_params` version-gated: main branch wraps `get_impl_cls()`/`update_graph_params` in `with set_current_vllm_config(vllm_config):` | | `vllm_ascend/models/deepseek_mtp.py` | [vllm#53106](vllm-project/vllm#53106) — removed `skip_prefixes` kwarg from `AutoWeightsLoader.__init__` | `AscendGlmMoeDsaForCausalLM.load_weights` version-gated: 0.27.1 uses `skip_prefixes=["rot."]`; main uses `WeightsMapper(orig_to_new_prefix={"rot.": None})` passed via `load_weights(weights, mapper=mapper)` | | `vllm_ascend/spec_decode/llm_base_proposer.py` | [vllm#52861](vllm-project/vllm#52861) — added `DeepseekV32MTPModel` to MTP architecture set in `model_returns_tuple()` | `model_returns_tuple` version-gated: 0.27.1 checks `{"DeepSeekMTPModel", "KimiK3MTPModel"}`; main also includes `"DeepseekV32MTPModel"` | | `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` | [vllm#52188](vllm-project/vllm#52188) — added `cp_rank`, `CP_SIZE`, `CP_INTERLEAVE` params to `_prepare_dflash_inputs_kernel` for DCP support | Entire `_prepare_dflash_inputs_kernel_ascend` kernel duplicated under `vllm_version_is("0.27.1")` gate: 0.27.1 uses 30 pos + 3 constexpr (no DCP params); main uses 31 pos + 5 constexpr | ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@cdc4824 --------- Signed-off-by: hfadzxy <starmoon_zhang@163.com>
### What this PR does / why we need it? The PR adapts vllm-ascend for compatibility with the latest vLLM main (commit `ba07e4a4`). | Files | Upstream vLLM change | vllm-ascend adaptation | |-------|---------------------|------------------------| | `.github/vllm-main-verified.commit` | — | Updated verified main commit hash from `cdc4824a21` to `ba07e4a48` | | `tests/e2e/conftest.py` | [vllm#53272](vllm-project/vllm#53272) — upstream plans to remove native Hunyuan V1/VL; [vllm#51665](vllm-project/vllm#51665) — dropped HunYuanVL `lm_head` workaround | Added `skip` condition for HunyuanVL e2e when `vllm_version_is("0.27.1")` is False (vLLM main) | | `tests/ut/core/test_profiling_chunk.py` | — vLLM main `Scheduler.__init__` reads `model_config.uses_mrope`, which infinitely recurses on a bare MagicMock | Version-gated: sets `type(model_config).uses_mrope = PropertyMock(return_value=False)` on main | | `tests/ut/core/test_recompute_scheduler.py` | — vLLM main `Scheduler.add_request` reads `spec_decode_metrics_level` | Version-gated: sets `scheduler.spec_decode_metrics_level = "none"` on main | | `tests/ut/patch/platform/test_patch_structured_output.py` | — Upstream changed structured output validation error type from `ValueError` to `VLLMValidationError` | Version-gated: `error_type = ValueError if vllm_version_is("0.27.1") else VLLMValidationError` used in all three fake validation functions and `pytest.raises` | | `vllm_ascend/attention/attention_v1.py` | [vllm#52839](vllm-project/vllm#52839) — moved `pcp.py` from `vllm.model_executor.layers.attention.pcp` to `vllm.v1.attention.ops.pcp` | Version-gated `_gather_prefill_cache_inputs` import path | | `vllm_ascend/compilation/acl_graph.py` | [vllm#49134](vllm-project/vllm#49134) — `get_current_vllm_config()` now raises `AssertionError` when called outside `set_current_vllm_config()` context | `update_full_graph_params` version-gated: main branch wraps `get_impl_cls()`/`update_graph_params` in `with set_current_vllm_config(vllm_config):` | | `vllm_ascend/models/deepseek_mtp.py` | [vllm#53106](vllm-project/vllm#53106) — removed `skip_prefixes` kwarg from `AutoWeightsLoader.__init__` | `AscendGlmMoeDsaForCausalLM.load_weights` version-gated: 0.27.1 uses `skip_prefixes=["rot."]`; main uses `WeightsMapper(orig_to_new_prefix={"rot.": None})` passed via `load_weights(weights, mapper=mapper)` | | `vllm_ascend/spec_decode/llm_base_proposer.py` | [vllm#52861](vllm-project/vllm#52861) — added `DeepseekV32MTPModel` to MTP architecture set in `model_returns_tuple()` | `model_returns_tuple` version-gated: 0.27.1 checks `{"DeepSeekMTPModel", "KimiK3MTPModel"}`; main also includes `"DeepseekV32MTPModel"` | | `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` | [vllm#52188](vllm-project/vllm#52188) — added `cp_rank`, `CP_SIZE`, `CP_INTERLEAVE` params to `_prepare_dflash_inputs_kernel` for DCP support | Entire `_prepare_dflash_inputs_kernel_ascend` kernel duplicated under `vllm_version_is("0.27.1")` gate: 0.27.1 uses 30 pos + 3 constexpr (no DCP params); main uses 31 pos + 5 constexpr | ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@cdc4824 --------- Signed-off-by: hfadzxy <starmoon_zhang@163.com>
… cudagraph vLLM vllm-project#52861 routed the DSA architectures onto the non-compiled V2 model runner / breakable-cudagraph path, but GlmMoeDsaForCausalLM was left out of the ROCm carve-out, flipping GLM-5.2 to MRV2 + breakable cudagraphs and regressing batch-1 decode TPOT by ~30-37% on gfx950 (FP8 and MXFP4). Add GlmMoeDsaForCausalLM to ROCM_DEFAULT_MRV1_ARCHITECTURES so GLM-5.2 stays on the compiled MRV1 path, and default breakable cudagraphs off entirely on ROCm (they regress performance today); VLLM_USE_BREAKABLE_CUDAGRAPH=1 still forces them on. Signed-off-by: Rohan Potdar <rohan.potdar@amd.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
### What this PR does / why we need it? The PR adapts vllm-ascend for compatibility with the latest vLLM main (commit `ba07e4a4`). | Files | Upstream vLLM change | vllm-ascend adaptation | |-------|---------------------|------------------------| | `.github/vllm-main-verified.commit` | — | Updated verified main commit hash from `cdc4824a21` to `ba07e4a48` | | `tests/e2e/conftest.py` | [vllm#53272](vllm-project/vllm#53272) — upstream plans to remove native Hunyuan V1/VL; [vllm#51665](vllm-project/vllm#51665) — dropped HunYuanVL `lm_head` workaround | Added `skip` condition for HunyuanVL e2e when `vllm_version_is("0.27.1")` is False (vLLM main) | | `tests/ut/core/test_profiling_chunk.py` | — vLLM main `Scheduler.__init__` reads `model_config.uses_mrope`, which infinitely recurses on a bare MagicMock | Version-gated: sets `type(model_config).uses_mrope = PropertyMock(return_value=False)` on main | | `tests/ut/core/test_recompute_scheduler.py` | — vLLM main `Scheduler.add_request` reads `spec_decode_metrics_level` | Version-gated: sets `scheduler.spec_decode_metrics_level = "none"` on main | | `tests/ut/patch/platform/test_patch_structured_output.py` | — Upstream changed structured output validation error type from `ValueError` to `VLLMValidationError` | Version-gated: `error_type = ValueError if vllm_version_is("0.27.1") else VLLMValidationError` used in all three fake validation functions and `pytest.raises` | | `vllm_ascend/attention/attention_v1.py` | [vllm#52839](vllm-project/vllm#52839) — moved `pcp.py` from `vllm.model_executor.layers.attention.pcp` to `vllm.v1.attention.ops.pcp` | Version-gated `_gather_prefill_cache_inputs` import path | | `vllm_ascend/compilation/acl_graph.py` | [vllm#49134](vllm-project/vllm#49134) — `get_current_vllm_config()` now raises `AssertionError` when called outside `set_current_vllm_config()` context | `update_full_graph_params` version-gated: main branch wraps `get_impl_cls()`/`update_graph_params` in `with set_current_vllm_config(vllm_config):` | | `vllm_ascend/models/deepseek_mtp.py` | [vllm#53106](vllm-project/vllm#53106) — removed `skip_prefixes` kwarg from `AutoWeightsLoader.__init__` | `AscendGlmMoeDsaForCausalLM.load_weights` version-gated: 0.27.1 uses `skip_prefixes=["rot."]`; main uses `WeightsMapper(orig_to_new_prefix={"rot.": None})` passed via `load_weights(weights, mapper=mapper)` | | `vllm_ascend/spec_decode/llm_base_proposer.py` | [vllm#52861](vllm-project/vllm#52861) — added `DeepseekV32MTPModel` to MTP architecture set in `model_returns_tuple()` | `model_returns_tuple` version-gated: 0.27.1 checks `{"DeepSeekMTPModel", "KimiK3MTPModel"}`; main also includes `"DeepseekV32MTPModel"` | | `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` | [vllm#52188](vllm-project/vllm#52188) — added `cp_rank`, `CP_SIZE`, `CP_INTERLEAVE` params to `_prepare_dflash_inputs_kernel` for DCP support | Entire `_prepare_dflash_inputs_kernel_ascend` kernel duplicated under `vllm_version_is("0.27.1")` gate: 0.27.1 uses 30 pos + 3 constexpr (no DCP params); main uses 31 pos + 5 constexpr | ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@cdc4824 --------- Signed-off-by: hfadzxy <starmoon_zhang@163.com>
vllm-project#52861 routed the DSA models to a fused norm+rope Triton kernel that writes the MLA KV cache itself and only supports fp8_ds_mla. vllm-project#51724 added nvfp4_ds_mla after that and the rebase missed it. Teach the fused kernel the nvfp4_ds_mla layout. Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
Purpose
Route
DeepseekV32ForCausalLM,GlmMoeDsaForCausalLM, and their MTP draftmodel through the CUDA
vllm.models.deepseek_v32implementation on everyNVIDIA GPU, while retaining the existing defaults elsewhere.
The optimized NVIDIA classes are currently unreachable from the model registry.
That leaves DeepSeek V3.2 and GLM-5.2 on the generic runner/graph path and
prevents their MTP draft model from using the matching implementation.
This PR:
NVIDIA GPU;
DeepseekV32ForCausalLMandDeepseekV32MTPclass namesthrough the package platform entry point;
the explicit AMD modules available for opt-in use;
fallbacks on earlier NVIDIA GPUs;
DeepseekV32MTPModelwith the main model;graphs on NVIDIA, while retaining the ROCm compiled default;
CompilationMode.NONEfor that graph path and preserves the explicitVLLM_USE_BREAKABLE_CUDAGRAPH=0opt-out; andThis supersedes #49790. That PR was automatically closed after its contributor
branch was synchronized to
main; it now has no commits or changed files andcannot be recovered by a maintainer push. The prepared changes were rebased onto
current
mainand published here instead.Duplicate-work check
No open PR covers this NVIDIA CUDA default-routing scope. In particular, #51915 is
an opt-in ROCm/MXFP4 correctness path and explicitly leaves the default registry
route unchanged. The other open GLM-5.2/DSA results are backend- or
kernel-specific. This remains the NVIDIA DSA routing item tracked by #48597.
Test Plan
Run the affected configuration tests:
Check that the routed target and draft model classes import through the registry:
.venv/bin/python -m pytest \ tests/models/test_registry.py::test_registry_imports \ -k 'DeepseekV32ForCausalLM or GlmMoeDsaForCausalLM or DeepseekV32MTPModel' -qRun pre-commit on all sixteen changed files.
Evaluate the full GSM8K set with TP=4 and MTP=3, then benchmark batch-size-1
serving on 4x NVIDIA GB200 with and without MTP. Both benchmark arms use the
same server command; the MTP arm adds the final
--speculative-configoption:CUDA_VISIBLE_DEVICES=0,1,2,3 .venv/bin/vllm serve \ nvidia/GLM-5.2-NVFP4 \ --served-model-name GLM-5.2 \ --revision aec724e8c7b8ee9db3b48c01c320f63f9cdaf8aa \ --tensor-parallel-size 4 \ --port 8300 \ --kv-cache-dtype fp8_e4m3 \ --max-model-len 16384 \ --max-num-seqs 256 \ --max-num-batched-tokens 16384 \ --no-enable-prefix-caching \ --gpu-memory-utilization 0.85 \ --safetensors-load-strategy prefetch \ --disable-uvicorn-access-log \ --kernel-config \ '{"ir_op_priority":{"rms_norm":["vllm_c","native"],"fused_add_rms_norm":["vllm_c","native"]},"enable_flashinfer_autotune":false}' \ --speculative-config '{"method":"mtp","num_speculative_tokens":3}'The serving environment has matching FlashInfer 0.6.17 packages:
flashinfer-python==0.6.17,flashinfer-cubin==0.6.17, andflashinfer-jit-cache==0.6.17+cu130.Test Result
The import/configuration checks cover CUDA-wide routing without requiring a
specific compute capability and confirm that the non-CUDA package entry point resolves
to the generic target and MTP classes. The full model evaluation and performance
runs below use GLM-5.2 on GB200; DeepSeek V3.2 and pre-SM100 GPU runtime were not
benchmarked here.
Full GSM8K, 1,319 questions, 5-shot, temperature 0, seed 42, max output 256,
concurrency 100:
Batch-size-1 SPEED-Bench
throughput_16k/low_entropy, 8,192 input tokens,1,024 output tokens, concurrency 1, one warmup and three measured requests:
1 / TPOT)MTP=3 improves wall-clock output throughput by 2.337x (+133.7%) and
decode-only throughput by 2.502x (+150.2%). Startup logs confirmed MRV2,
automatic breakable FULL + PIECEWISE CUDA graphs,
CompilationMode.NONE, sparseFlashInfer MLA,
FLASHINFER_TRTLLMNVFP4 MoE,allreduce_rmsfusion, andDeepseekV32MTPModelfor the MTP arm.A forced FlashMLA smoke and batch-size-1 benchmark also passed on 4x GB200 with the real NVFP4 checkpoint. The server used
--attention-backend FLASHMLA_SPARSE --kv-cache-dtype fp8_ds_mla, TP=4, V2 model runner, and breakable full + piecewise CUDA graphs. Logs confirmedFLASHMLA_SPARSEattention andFLASHINFER_TRTLLMNVFP4 MoE. A 2,048-input/1,024-output SPEED-Bench request achieved 114.01 output tok/s including TTFT, 8.75 ms TPOT (about 114.3 decode tok/s), and 34.48 ms TTFT. The request completed successfully with no runtime JIT warning.Compile-test cleanup:
DeepSeek V3.2 was removed from the H100 compile-startup suite and from the CUDA NVFP4 fusion-E2E suite because the NVIDIA path defaults to breakable CUDA graphs with compilation disabled. The generic ROCm/AITER compiler-pass tests remain because AMD intentionally stays on the compiled path.
AI assistance was used to prepare the rebase/default-path changes, debug the
benchmark environment, run tests and evaluations, and draft this description.
The human submitter must review every changed line and the evidence above and
own the change end-to-end before marking this draft ready for review.
Essential Elements of an Effective PR Description Checklist
BEFORE SUBMITTING, PLEASE READ https://docs.vllm.ai/en/latest/contributing